Papers with spoken language processing

5 papers
Sounding Board: A User-Centric and Content-Driven Social Chatbot (N18-5)

Copied to clipboard

Challenge: Sounding Board is a social chatbot that can hold a coherent conversation with humans . the system is user-centric in that users can control the topic of conversation, while the system adapts to the user's needs.
Approach: They present Sounding Board, a social chatbot that won the 2017 Amazon Alexa Prize.
Outcome: The system is user-centric in that users can control the topic of conversation, while the system adapts to the user's needs.
CPJD Corpus: Crowdsourced Parallel Speech Corpus of Japanese Dialects (L18-1)

Copied to clipboard

Challenge: Various corpora of dialects have been collected using a well-equipped recording environment due to geographical and expense issues.
Approach: They construct a crowdsourced parallel speech corpus of Japanese dialects using crowdsourcing platforms.
Outcome: The proposed corpus includes parallel text and speech data of 21 Japanese dialects.
MYCanCor: A Video Corpus of spoken Malaysian Cantonese (L18-1)

Copied to clipboard

Challenge: The corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers.
Approach: the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers.
Outcome: the corpus consists of 20 hours of video recordings of spontaneous talk-in-interaction typically involving 2-4 speakers.
SMASH Corpus: A Spontaneous Speech Corpus Recording Third-person Audio Commentaries on Gameplay (2020.lrec-1)

Copied to clipboard

Challenge: Developing a spontaneous speech corpus is important for spoken language research . a corpus of spontaneous speech is needed to develop these techniques .
Approach: They propose to use Japanese male commentators' spontaneous speech to construct a SMASH corpus . they use transcriptions and topic tags to annotate the commentaries and report some results .
Outcome: The proposed corpus includes spontaneous speech of two Japanese male commentators . the authors report that the annotations yielded a better corpus than the previous methods .
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model (2026.acl-long)

Copied to clipboard

Challenge: Phone-level modeling of speech is a common approach to speech recognition, but it relies on task-specific architectures and datasets.
Approach: They propose a phonetic framework capable of performing multiple phone-related tasks . they propose 'Phonetic Open Whisper-style Speech Model' that can perform these tasks together .
Outcome: The proposed model outperforms or matches specialized PR models of similar size while supporting G2P, P2G, and ASR.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations